
download llamafile from github repo : https://github.com/mozilla-ai/llamafile/releases/tag/0.10.6
rename it to .exe
if using flash format it as extFAT
go to Hugging Face and downoad GGUF format of LLMS you want
open notepad and create a file named start.bat:
./llafial*.*.exe --server --model "llm-gguf-file.gguf" --ctx-size 32768
./llafial*.*.exe --server --embeddings --model "llm-gguf-file.gguf" --ctx-size 32768

--host 0.0.0.0 
--port 8081

sc.exe create XLLMEmbdService binPath= "D:\Projects\LLMs\bge-m3.bat"
sc.exe create XLLMService binPath= "D:\Projects\LLMs\gemma3b1sc.bat"

sc queryex type=service state=all
sc queryex type=service state=all | find /i "SERVICE_NAME:"
sc queryex type=service state=active

sc query XLLMService
sc query XLLMEmbdService

sc delete XLLMService
sc delete XLLMEmbdService

https://docs.mozilla.ai/llamafile/using-llamafile/api
https://github.com/mozilla-ai/llamafile

-------------------

Uncesored Qwen 3.5 Coder GGUF:

https://huggingface.co/DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF

-------------------

Download these files and put them inside xAiApi/Models:

https://huggingface.co/ggerganov/whisper.cpp/blob/main/ggml-base.bin
https://huggingface.co/ggerganov/whisper.cpp/blob/main/ggml-small.bin